Back

BMC Medical Genomics

Springer Science and Business Media LLC

Preprints posted in the last 90 days, ranked by how well they match BMC Medical Genomics's content profile, based on 50 papers previously published here. The average preprint has a 0.04% match score for this journal, so anything above that is already an above-average fit.

1
FIND: a software tool for identifying population-enriched pathogenic variants in gnomAD

Horowitz, A. L.; Liebman, A. Z.; Liebman, S. W.

2026-06-10 genomics 10.64898/2026.06.05.730273 medRxiv
Top 0.1%
8.7%
Show abstract

Founder mutations are variants that arose in a single ancestor and became enriched in a descendant population through a bottleneck and endogamy. Identification of pathogenic founder mutations has facilitated efficient targeted screening. More broadly, even without confirmed founder status, identifying pathogenic variants that are enriched within specific populations reveals population-specific disease burden. However, many such variants remain hidden in plain sight within existing datasets. To address this gap, we developed FIND (Founder candidates hidden IN Data), a web tool that identifies pathogenic, likely pathogenic, and predicted loss-of-function variants in gnomAD with frequencies >0.00008 in one ancestry group and at least tenfold higher than in all others (after zeroing populations with four or fewer observed alleles). Testing FIND on the genes FLNC, TMEM127, MYH7, and BRCA2 confirmed its utility and functionality by identifying nine well-known founder mutations and seven candidate founders. Candidates enriched in African American and admixed American populations were validated with the All of Us database, highlighting the utility of this approach for populations historically underrepresented in genetic studies. Source code is freely available at https://github.com/aacoder105/FIND under an MIT license, with a web interface at https://ethnic-variant-mutation-finder.onrender.com/.

2
Evaluating Aggregated Gene Level eQTL Scores

Meyer, D.; Popko, N.; Laub, D.; Schofield, P.; Amariuta, T.; Alexandrov, L. B.; Carter, H.

2026-08-26 bioinformatics 10.64898/2026.08.21.746287 medRxiv
Top 0.1%
7.9%
Show abstract

Genetic feature engineering, used in methods such as transcriptome-wide association study, supports gene-trait association testing by aggregating single variants into gene-level features predictive of expression. To evaluate how different model architectures, LD filtering thresholds, and variant prioritization methods affect expression prediction quality, we trained over 3 million models and evaluated their performance in independent cohorts. Using the best performing models to impute expression and immunotherapy response as an example trait, we found a significant association with the reactive oxygen species pathway (p=0.032). Our model training workflow will support genetic feature engineering towards improved complex trait modeling.

3
Uncovering High-Order Epistatic Interactions in GWAS via a Machine Learning-Based Feature Engineering Framework

Byun, J.; Saha, D.; Han, Y.; Shaw, V. R.; Siminovitch, K.; Amos, C. I.

2026-08-09 genomics 10.64898/2026.08.03.742638 medRxiv
Top 0.1%
6.7%
Show abstract

BackgroundGenome-wide association studies (GWAS) often fail to identify higher-order epistatic interactions that contribute to complex inheritance patterns of traits and diseases. While machine learning (ML) can capture non-linear relationships, extracting interpretable insights from these models remains a challenge. We propose a novel tree-based feature engineering framework that uses Classification and Regression Trees (CART) to explicitly encode high-order interaction decision paths as dummy variables. We investigate three path-based encoding strategies: (i) all decision paths, (ii) leaf-node paths only, and (iii) internal-node paths only. This approach aims to transform complex decision boundaries into discrete features that capture nonlinear interactions that are not readily captured by traditional association models. ResultsThe framework was evaluated using genetic data for ANCA-associated vasculitis (AAV). To manage the high dimensionality of the engineered feature space, we applied a comprehensive suite of ML methods across three tasks: (1) Ensemble Learning (Random Forest, XGBoost, and Gradient Boosting Machine); (2) Decision Tree Analysis (CART); and (3) Regression and Classification Tasks (Regularized Linear Regression/LASSO, Support Vector Machine, and Logistic Regression). Stepwise feature selection and regularization were employed to isolate the most informative interaction patterns. Results indicate that incorporating CART-derived interaction paths--particularly those from high-impact regions of the tree--significantly improves classification accuracy and model interpretability compared to using the original feature space alone. ConclusionsThe proposed framework provides a robust, scalable methodology for identifying high-order genetic interactions. By bridging the gap between the predictive power of ensemble ML and the necessity for mechanistic insight, this approach offers a clearer mapping of the combinatorial genetic processes underlying complex diseases. While applied here to AAV, the method is highly adaptable for exploring the genetic architecture of diverse populations and complex traits.

4
Systematic benchmarking of low-input whole exome sequencing workflows for longitudinal ctDNA profiling in pancreatic ductal adenocarcinoma

James, L. G.; Thorn, G. J.; Morel, C.; PCRFTB, ; Kocher, H. M.; Ross-Adams, H. E.; Chelala, C.

2026-07-03 genomics 10.64898/2026.06.29.734743 medRxiv
Top 0.1%
6.3%
Show abstract

Whole exome sequencing (WES) of circulating tumour DNA (ctDNA) enables longitudinal monitoring of tumour dynamics, evolution and treatment response but remains technically challenging in low-input, low-shedding settings such as pancreatic ductal adenocarcinoma (PDAC). Here, we systematically compared three commercially available low-input WES workflows incorporating Agilent (V6, V8) and Qiagen exome capture designs using ultra-low input cfDNAs extracted from multiple matched longitudinal plasma samples from PDAC patients. Using predefined performance metrics including coverage, duplication rate and variant detection and additional metrics relevant for clinical genomic profiling in patient care, we show that all three workflows produced high-quality sequencing data, even from very low input cfDNA. Within the conditions tested here, the Agilent V8 workflow provided the most favourable balance of coverage uniformity, sequencing efficiency and hotspot coverage for low input, low tumour fraction cfDNA WES. These findings demonstrate that workflow design, including capture footprint, substantially influences ctDNA WES performance in low-input clinical contexts. These findings are particularly relevant in early stage and/or minimal residual disease settings, where tumour fractions are low and recovery of genomic information from limited-input samples is critical.

5
Integration of lung tissue proteomics and genome-wide association data to identify lung cancer susceptibility proteins and potential drug targets

Xu, S.; Shi, J.; Shu, X.-O.; Tao, R.; Dou, Y.; Guo, X.; Wen, W.; Yang, Y.; Zhang, B.; Wu, J.; Deppen, S. A.; Li, B.; Zheng, W.; Long, J.; Cai, Q.

2026-06-22 epidemiology 10.64898/2026.06.18.26355973 medRxiv
Top 0.1%
5.5%
Show abstract

Background: Proteins directly impact disease development and act as drug targets. Therefore, we integrated genomic and lung tissue proteomics data to identify lung cancer susceptibility proteins, elucidating genetic mechanisms and candidate drug targets. Method: We profiled the proteome and genome in non-neoplastic lung tissue from 200 lung cancer patients. Using this data, we constructed genetic models to predict abundance across the proteome in lung tissue. We applied these models to genome-wide association study (GWAS) data from 55,174 lung cancer cases and 1,294,174 controls to evaluate their associations with the risk of lung cancer, overall and by major histological subtypes. Bayesian colocalization and Mendelian randomization (MR) analyses were used to prioritize putative causal proteins, which were cross-referenced with three main drug-protein databases to identify potential therapeutic targets. Results: We identified 29 proteins associated with lung cancer risk at a false discovery rate < 5%, including 25 for overall lung cancer, two (AQP3 and IL18) specifically for adenocarcinoma, and another two (HMGN2 and HLA-DMB) for squamous cell carcinoma. Of them, genes encoding 17 proteins reside at least 2Mb away from any known GWAS risk loci, including 14 for overall lung cancer (HYI, GPX1, GMPPB, DSP, HDDC2, MTCH2, SUOX, JMJD7, PDIA3, IL16, IQGAP1, SULT1A2, ARHGAP27, and TYMP) and three for subtypes (AQP3, IL18, and HMGN2). Among the 12 proteins located within the known risk loci, EPHX2, CLDN18, PSMD5, and CYP2S1 proteins showed an association independent of the proximal GWAS-identified lead variant. Colocalization and/or MR analysis suggested 11 potential causal proteins. Five of these candidate causal proteins (DSP, CLDN18, IQGAP1, IL18 and TYMP) are targeted by nine drugs already approved by the FDA or in phase III trials. Conclusion: Our study identified novel lung cancer susceptibility proteins and potential drug targets, offering valuable insights into lung cancer biology and future translational utilities.

6
A Custom Global Screening Array for Integrated Familial Hypercholesterolemia Detection and Polygenic Risk Assessment in a Multi-Ethnic New Zealand Population

Vikhorev, A.; Struchalin, M.; Sun, X.; Wen, Y.; Wihongi, H.; Gladding, P.

2026-06-24 cardiovascular medicine 10.64898/2026.06.22.26355820 medRxiv
Top 0.1%
5.5%
Show abstract

Background: Cardiovascular disease (CVD) is the leading cause of mortality in New Zealand, with significant inequities affecting M[a]ori and Pacific peoples. Familial hypercholesterolaemia (FH) affects approximately 1 in 313 individuals globally, yet over 90% remain undiagnosed. Standard polygenic risk scores (PRS) derived from European cohorts may not be portable to diverse ancestries. We developed the HoloQ Omniscan Waka Te Ira, a custom Illumina Global Screening Array (GSA) v3 enriched with FH mutations, coronary artery disease (CAD) PRS markers, and network medicine-derived content. Methods: We customised the GSA v3 by adding 43,437 single nucleotide polymorphisms (SNPs) targeting FH and CAD. Content included 6,717 unique variants in primary FH genes; 14,005 pathogenic or likely pathogenic cardiovascular and pharmacogene variants; and 5,845 copy number variant probes. We further incorporated 5,232 network medicine derived CAD SNPs, 14,806 rare variants for a multiancestry PRS, and 407 globally diverse and population-specific variants. The final design comprised 47,027 target SNPs. Validation utilised large-scale genotype and whole-genome sequencing (WGS) datasets with PRS benchmarking. Results: In a large European-ancestry dataset, we observed high recovery for common PRS loci but low recovery for population-specific founder variants. The array captured 938 (84%) of all pathogenic or likely pathogenic FH variants catalogued in ClinVar, representing a 26.4% expansion beyond the standard backbone array. WGS validation identified additional carriers of rare high impact variants present only in the custom content. The selected CAD PRS model achieved an adjusted area under the receiver operating characteristic curve of 0.786. Conclusion: The HoloQ Omniscan Waka Te Ira enhances detection of clinically relevant FH variants and provides robust PRS coverage. The low recovery of population-specific alleles underscores the necessity of this custom array for equitable genomic medicine in New Zealand's multi-ethnic population.

7
From Data Curation to Risk Reporting: A Pipeline for Polygenic Risk Scores

Barbosa Araujo, P. V.; da Silva Fiuza, T.; Ferraz, R. S.; Kroll, J. E.; Andrade, R. L.; Gomes, D. H. F.; Varuzza, L.; de Souza, G. A.; de Souza, S. J.

2026-08-20 bioinformatics 10.64898/2026.08.12.743994 medRxiv
Top 0.1%
5.3%
Show abstract

Polygenic risk scores (PRS) have emerged as a powerful tool for quantifying genetic susceptibility to complex traits and diseases. However, their calculation and interpretation require standardized data curation, robust statistical methods, and clear reporting strategies. In this work, we present an integrated pipeline designed to address these challenges. The pipeline begins with the construction of a curated genotype/phenotype database derived from public repositories, ensuring that only phenotypes with appropriate metadata, statistical distributions, and ethical suitability are retained. The final dataset comprises 2,346 phenotypes covering 38,256,468 unique SNPs. These phenotypes serve as the final analytical units for PRS calculation, risk stratification, and individual-level interpretation. The generated reports integrate sample-level results, phenotype categorization, risk classification, study references, and variant tables, providing a structured and interpretable output for end users. Together, the curated database and reporting framework establish a comprehensive toolbox for PRS analysis, enhancing reproducibility, transparency, and usability in both research and clinical contexts.

8
Using somatic data to aid germline clinical variant interpretation in developmental disorders

Andrews, K. A.; Neville, M. D.; Martincorena, I.; Rahbari, R.; Firth, H.; Lindsay, S. J.; Tischkowitz, M.; Hurles, M.

2026-06-21 genomics 10.64898/2026.06.17.732808 medRxiv
Top 0.1%
5.3%
Show abstract

Accurate interpretation of rare germline variants remains a major challenge in developmental disorders (DD). Somatic mutation data represent a largely untapped source of evidence for germline variant classi-fication. Identical or nearby mutations that drive positive selection when present in somatic tissues can cause developmental disorders when present in the germline. We integrated somatic mutation data from the Catalogue Of Somatic Mutations In Cancer (COSMIC), and healthy tissues (sperm and buccal epithelium) with germline variant datasets from ClinVar and large studies of de novo mutations in DD patients. Across 970 dominant DD genes, 195 have evidence of somatic selection, with a majority demonstrating concordant mechanisms between germline and somatic contexts. We benchmark the ability of somatic data to discriminate pathogenic from benign germline missense variation across dominant DD genes, identifying 145 genes in which somatic data are informative. The strongest utility is in altered-function genes where germline and somatic mechanisms are concordant, for example the RASopathy genes. In these genes, codon-level aggregation of somatic missense counts yields predictive performance comparable to computational predictors or MAVE assays (AUC-ROC 0.895 for somatic data, versus 0.893 for REVEL). Combining somatic features with computational scores improves discrimination further. Using likelihood ratios, we map COSMIC missense codon count thresholds onto American College of Medical Genetics and Genomics/Association for Molecular Pathology (ACMG/AMP)-style evidence strengths, showing that somatic data can reach strong levels of evidence in germline variant interpretation in DD and enable reclassification of variants of uncertain significance. Together, these results establish somatic mutation data as a scalable and clinically actionable evidence source for germline variant interpretation in select DD genes. Graphical abstract(Generated using FigureLabs) O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=104 SRC="FIGDIR/small/732808v1_ufig1.gif" ALT="Figure 1"> View larger version (40K): org.highwire.dtl.DTLVardef@150bec9org.highwire.dtl.DTLVardef@1dacf5org.highwire.dtl.DTLVardef@46121dorg.highwire.dtl.DTLVardef@4f5c38_HPS_FORMAT_FIGEXP M_FIG C_FIG

9
Droplet Digital PCR as a First-Line Detection Tool in the Genetic Diagnosis of Vascular Anomalies

Lane, T.; Green, T. E.; Garza, D.; Brown, N. J.; de Silva, M. G.; Bennett, M. F.; Tubb, C.; Macdonald, S. M. W.; Gascoigne, A.; Phillips, R. J.; Slavin, J.; D'Arcy, C.; MacGregor, D.; Clifford, A.; Pathmanathan, L.; Robertson, S. J.; Bekhor, P.; Simpson, J.; Gooley, S.; Scheffer, I. E.; Berkovic, S. F.; Penington, A. J.; Hildebrand, M.

2026-08-14 genetic and genomic medicine 10.64898/2026.08.11.26359368 medRxiv
Top 0.1%
5.1%
Show abstract

Targeted precision therapies are increasingly used in the treatment of individuals with vascular anomalies (VAs). This increases the need for rapid, accurate and inexpensive genetic diagnosis. Droplet digital polymerase chain reaction (ddPCR) is an alternative to next-generation sequencing (NGS), permitting rapid, highly sensitive interrogation of recurrent pathogenic mosaic variants. We examined the feasibility of ddPCR as a primary diagnostic tool in a large cohort of individuals with VAs. Lesional tissue was collected for ddPCR of up to 46 recurrent pathogenic variants across 16 genes associated with VAs. Specimens were assessed on a subset of assays for each individual based on clinical phenotype. Most individuals who had negative ddPCR results went on to high-depth gene panel or deep exome NGS, or Sanger sequencing. Here we report the phenotypic and molecular findings for 78 newly recruited and tested individuals in addition to the 60 individuals already reported from our cohort. The overall diagnostic yield for our cohort when combined with individuals previously reported was 104/138 (75%). Of 138 individuals tested, recurrent pathogenic variants were detected in 71 (51%) on ddPCR. Variants were most frequently identified in PIK3CA (n=28), TEK (n=18), GNAQ (n=12), or MAP2K1 (n=7). In a further 33 individuals, pathogenic variants were identified on NGS or Sanger sequencing. Our findings indicate that ddPCR is an efficient method achieving a high diagnostic yield in our cohort when used prior to sequencing.

10
Uveal and cutaneous melanoma share a common mutation with distinct prognostic implications: A bioinformatic study

Razmjooei, F.; Ashayeri, H.; Jafarzadeh, Z.; Dabbaghabdollahi, P.; Jafarizadeh, A.

2026-08-11 genetic and genomic medicine 10.64898/2026.08.07.26359988 medRxiv
Top 0.1%
5.1%
Show abstract

Background: Uveal melanoma (UM) and cutaneous melanoma (CM) both originate from the same cell line. This proposes the possibility of a shared mechanism between entities, requiring explicit investigation. Methods: Data from GWAS Catalog and DisGeNET were used to identify shared variation-disease associations (VDAs) between UM and CM. The results were validated using the Ensembl database. In the next step, the STRING database was used to identify the protein-protein interaction. Results: Subsequently, 109 unique VDAs were identified for UM and 880 for CM. However, only 2 VDAs were found to be shared among UM and CM in different ethnic groups. These shared VDAs were rs12203592 of the IRF4 gene, rs12913832 of the HECT and RLD domain-containing E3 ubiquitin protein ligase 2 (HERC2) gene. Notably, PPI network assessment through STRING showcased that OCA2 and IRF4 directly interacted with HERC2. Conclusion: While HERC2 acts as a poor prognostic factor in uveal melanoma, IRF4 status is a key prognostic indicator in both UM and CM. Identifying IRF4 allele contributions enables a better understanding of melanoma pathogenesis and fosters the development of disease-specific approaches.

11
Asthma Exacerbations: Integrative Analysis of miRNA Activity Using Single-Cell Transcriptomics

Hadikhani, P.; Yan, X.; Chupp, G. L.; Ban, G. Y.; Piparia, S.; McGeachie, M.; Sharma, R.; Weiss, S. T.; Laurent, L. C.; Kho, A. T.; Tantisira, K. G.

2026-08-06 bioinformatics 10.64898/2026.07.31.741637 medRxiv
Top 0.1%
4.9%
Show abstract

BackgroundAsthma exacerbations are caused by dysregulated cellular interactions between airway and immune cell populations. Circulating microRNAs (miRNAs) are potential biomarkers for asthma exacerbations; however, their target airway cells remain poorly defined. ObjectiveTo identify the cell types that are regulated by the circulating microRNAs linked to asthma exacerbations and the extent to which the cells are regulated by miRNAs. MethodsWe integrated a curated panel of exacerbation-associated circulating miRNAs with single-cell RNA sequencing (scRNA-seq) profiles from induced sputum of 16 asthma patients and 8 healthy controls. Experimentally validated miRNA-target interactions were combined with cell-type-specific differential expression. Elastic Net regression and SHAP analysis quantified gene-level regulatory contributions, yielding a composite Regulation Strength metric. Findings were validated against four independent GEO datasets. ResultsImmune cells, including monocytes, dendritic cells, and macrophages, demonstrated the strongest statistically significant miRNA regulatory signals, in contrast to airway epithelial cells.hsa-miR-222-3p showed opposing regulatory effects in mature versus alveolar macrophages, indicating differentiation-state-dependent activity, while B_Plasma cells showed no detectable regulatory effect from any miRNA tested. Independent GEO validation confirmed higher expression of protective miRNAs (hsa-miR-126-3p, hsa-miR-146b-5p) in healthy individuals, consistent with prior CAMP cohort associations. ConclusionCirculating miRNAs show cell-type-specific regulatory activity, strongest in monocytes, dendritic cells, and macrophages. hsa-miR-222-3p showed opposing regulatory directions between macrophage subtypes, while B_Plasma cells showed no effect, validated across independent GEO cohorts.

12
Integrative optical genome mapping and long-read sequencing resolve constitutional complex rearrangements at nucleotide resolution

Burssed, B.; van der Sanden, B.; Hops, W.; Neveling, K.; Kamping, E.; van Beek, R.; den Ouden, A.; Derks, R.; Timmermans, R.; Perrone, E.; Ramos, M. A.; Bellucco, F. T.; Hoischen, A.; Melaragno, M. I.

2026-08-28 genomics 10.64898/2026.08.27.747510 medRxiv
Top 0.1%
4.8%
Show abstract

Complex rearrangements are one of the rarest types of structural variants (SVs) and can be divided into two categories: complex chromosomal rearrangements (CCRs) and complex genomic rearrangements (CGRs). CCRs include structural rearrangements that present at least three breakpoints and show exchange of genetic material between more than two chromosomes and CGRs are rearrangements that present more than one junction and/or more than one SV in cis. They are usually formed by one of the chromoanagenesis mechanisms, where a massive disruptive cellular event leads to multiple structural rearrangements. Classical cytogenomic techniques have been commonly applied for their characterization, but methodologies that involve longer DNA molecules, namely optical genome mapping (OGM) and long-read genome sequencing (lrGS), present a considerably higher SV detection resolution, revealing more details about the rearrangements, including precise breakpoint location. Here, we describe six patients with complex rearrangements investigated through a combination of different techniques: karyotyping, chromosomal microarray, and OGM were performed to characterize the rearrangements. Subsequently, lrGS was used to further resolve the alterations, refine their breakpoints' location, and sequence their junction points. Three patients presented CCRs involving three, four, and six chromosomes, while three exhibited CGRs involving one different chromosome each, providing a variety of complex SVs to show the importance of each technique and their combination in rearrangement resolution. In total, the complex rearrangements presented 127 breakpoints, 66 junction points and involved 14 of the 24 chromosomes. Higher-resolution techniques revealed additional complexity in all cases. Despite the advances provided by OGM and lrGS, conventional karyotyping remained indispensable for complete rearrangement resolution. In two patients, the findings supported a novel mechanism combining features of the different chromoanagenesis processes. Furthermore, evidence of inherited alterations was identified, and the comprehensive characterization of the rearrangements enabled more accurate genotype-phenotype correlations. Our findings indicate that an integrated approach combining karyotyping, OGM, and lrGS can completely resolve SVs, including complex rearrangements.

13
"Transcriptional and isoform-level regulation of lipid-candidate genes in preeclamptic placentas"

Eyer, K. S.; Lemaire, M.; Fan, X.; Wilson, S. L.

2026-08-21 genomics 10.64898/2026.08.17.745256 medRxiv
Top 0.1%
4.3%
Show abstract

Preeclampsia (PE) is a hypertensive pregnancy-specific disorder and a leading cause of maternal and fetal mortality. A common feature of PE placentas and maternal plasma is dyslipidemia, or abnormal lipid levels, which can increase oxidative stress and endothelial dysfunction. However, the precise transcriptional, post-transcriptional, and epigenetic mechanisms underlying these abnormalities remain poorly characterized. Identifying such changes may clarify disease mechanisms and identify lipid-related PE biomarkers. We conducted a large-scale meta-analysis integrating public placental datasets from NCBI GEO, comprising four DNA methylation (DNAm) datasets (n = 172), three RNA-sequencing datasets (n = 92), and an independent RNA microarray validation cohort (n =146). We evaluated differential DNAm (limma), gene expression (DESeq2), transcript-level shifts (Swish), and alternative splicing (rMATS) in PE versus control placentas, with all analyses stratified by fetal sex via an interaction term model. We also performed placental cell-type deconvolution to quantify PE-associated cell-type proportion changes. Our results demonstrated that lipid-related regulation changes in PE placentas occur primarily at the gene and transcript level, with DNAm showing no changes. We also identified significant isoform switching in PE that were undetected by differential gene expression analysis, and primarily driven by alternative transcription initiation and termination sites rather than alternative splicing. A subset of these isoform switches mapped to pathways dysregulated in PE and were predicted to cause functional protein changes. An interaction term model identified several sex-specific differentially expressed genes (DEGs) in PE, including a subset of male-specific downregulated genes involved in oxidative metabolism. However, many of the remaining sex-specific DEGs across both sexes were previously uncharacterized in the literature. These findings suggest that transcriptional and isoform-level regulation play a role in PE-associated dyslipidemia, with certain regulatory pathways displaying fetal sex-specific patterns. Highlights- Preeclampsia-associated dyslipidemia manifests at the gene and transcript level - Reciprocal isoform switches were missed by standard gene-level analyses - Alternative transcript initiation and termination drove isoform switching - Sex-interaction modeling identified sex-specific transcriptional shifts in PE

14
Early Tracheal and Salivary miRNAs in Extremely Preterm Infants Predict BPD-related Pulmonary Hypertension

Li, T.; Zhang, S.; Aluquin, V.; Donnelly, A.; Stephens, H.; Sharma, S.; Hicks, S. D.; Liu, D.; Austin, E.; Siddaiah, R.

2026-06-23 bioinformatics 10.64898/2026.06.17.732493 medRxiv
Top 0.1%
4.2%
Show abstract

Pulmonary hypertension (BPD-PH) associated with bronchopulmonary dysplasia (BPD) in preterm infants associates with high morbidity and mortality within the first two years of life. In a previous unbiased study, we identified a panel miRNAs in tracheal aspirates (TA) that were differentially expressed in extremely low gestational age newborns (ELGANs) with BPD-PH compared to those with BPD but no PH. To explore the predictive potential of these miRNAs, we studied TA exosomes from 7 days old ELGANs and analysed a curated panel of 16 miRNAs through logistic regression and calculated the predictive AUROC to diagnose BPD-PH at 36 weeks PMA. AUROC of TA miRNAs was 0.76 with sensitivity and specificity of 53% and 93%, respectively. Adding sex and gestational age to the variables improved the AUROC to 0.78 with sensitivity and specificity of 61 and 87% respectively. Due to challenges of obtaining TA in non-invasively ventilated infants, we collected saliva samples from ELGANs at 7 days of age and compared the log expression of these 16 miRNAs in both biofluids and found significant correlation in their expression (pearson r=0.92, p<0.001). We calculated the predictive AUROC of the same miRNAs to diagnose BPD-PH at 36 weeks PMA. AUROC of these miRNAs in saliva was = 0.85 with sensitivity and specificity of 82% and 72%, respectively; addition of biological sex and gestational age improved AUROC to 0.86 with sensitivity and specificity of 79% and 76% respectively. Leave-one-sample-out sensitivity analysis demonstrated stable training performance with reduced performance in testing samples, supporting the need for validation in larger independent cohorts. In conclusion, early salivary miRNAs have great potential for risk stratification of ELGANs to develop BPD-PH, while also providing the opportunity to identify target molecules and mechanisms that modulate molecular function.

15
Machine Learning-based Prediction of Preterm Birth Using Genetic Data

Sundelin, H.; Jacobsson, B.; Ytterberg, K.; Sole-Navais, P.; Juodakis, J.

2026-06-26 genetic and genomic medicine 10.64898/2026.06.24.26356330 medRxiv
Top 0.1%
4.1%
Show abstract

The leading cause of mortality and morbidity in children under the age of 5 is preterm birth. The timing of birth is influenced by both genetic and environmental factors, but the underlying mechanisms remain poorly understood, making its prediction difficult. In this study, we investigated the potential of using machine learning models to predict preterm birth based on genetic data from the Norwegian Mother, Father and Child Cohort Study (MoBa). We trained and evaluated several classification algorithms on individual-level genetic data from over 15,000 mothers and children. Our results indicate that the predictive capacity of maternal gestational duration-associated loci for preterm birth is limited, with the highest AUC values around 0.57. Additionally, incorporating more SNPs within the associated loci did not improve prediction performance. As expected, the contribution of the maternal genome to preterm birth prediction was found to be larger than that of the fetal genome. Overall, our findings suggest that while genetic testing provides some information about an individual's risk for preterm birth, further research incorporating additional factors is necessary to enhance predictability.

16
SPICE: A Robust Computational Framework for Identifying Copy Number Variations in Spatial Transcriptomics

Banerjee, K.; Langefeld, R. C.; Keller, E. T.; Zhou, X.

2026-07-04 genomics 10.64898/2026.06.30.735508 medRxiv
Top 0.1%
4.0%
Show abstract

Copy number variation (CNV), which alters the number of genomic segments, is a major driver of intratumor heterogeneity, characterized by spatially organized and genetically distinct cell populations. Recent advances in spatially resolved transcriptomic (SRT) technologies, which profile gene expression across thousands of spatially indexed tissue locations, offer a powerful opportunity to reconstruct the CNV architecture and dissect the spatial organization of cancer subclones. Here, we introduce SPICE (spatial inference of CNV events), a probabilistic method for identifying somatic CNVs and allele-specific copy number (ASCN) profiles from SRT data. A key feature of SPICE is its ability to integrate multiple complementary information available in SRT data, including gene expression, spatial coordinates, and heterozygous SNPs inferred from transcriptomic reads, to substantially enhance the accuracy and power of CNV detection. Using datasets generated across different SRT platforms, we first assess the reliability of SNPs derived from SRT data to ensure robust downstream inference. We then demonstrate that SPICE effectively integrates these modalities to deliver accurate and spatially coherent reconstruction of CNV landscapes and subclonal architecture, while maintaining excellent control of false discoveries. Together, SPICE provides a robust and effective solution for dissecting genomic heterogeneity in SRT studies of cancer.

17
ClinGen Glaucoma Variant Curation Expert Panel recommendations enhance classification of myocilin variants

Tompson, S. W. J.; Graham, P.; Hadler, J.; Pasutto, F.; Whisenhunt, K. N.; Chakrabarti, S.; Young, T. L.; Craig, J. E.; Hewitt, A. W.; Siggs, O. M.; Hulleman, J. D.; Mackey, D. A.; Burdon, K. P.; Dubowsky, A.; Souzeau, E.

2026-07-31 genetic and genomic medicine 10.64898/2026.07.29.26359282 medRxiv
Top 0.1%
4.0%
Show abstract

Pathogenic variants in the myocilin (MYOC) gene are the most common cause of Mendelian open-angle glaucoma. In 2022, the Clinical Genome Resource (ClinGen) Glaucoma Variant Curation Expert Panel (VCEP) published rule specifications for MYOC variant interpretation, including a pilot study of 81 variants. Here, we present the results of curating 271 MYOC variants reported in people with open-angle glaucoma using updated specification rules. Of all the variants, 11 were classified as benign (B), 45 as likely benign (LB), 166 as variants of uncertain significance (VUS), 35 as likely pathogenic (LP), and 14 as pathogenic (P). All LP/P variants were located within the conserved olfactomedin domain encoded by exon 3. The updated variant curation guidelines from the Glaucoma VCEP increased the number of clinically definitive classifications from 28% (74/265) to 39% (105/271), with 95% (41/43) of reclassified variants moving to greater clinical relevance. Functional evidence was lacking for 93% (154/166) of VUS. Additional functional evidence could further enhance classification by halving (85/166) the proportion of those classified as VUS. These findings highlight the role of rule calibration and rigorous functional evidence assessment toward improving variant classification with clinical utility for patients.

18
GLOF: A large-scale expert-curated benchmark dataset of gain-of-function and loss-of-function missense variants

Maricato, V.; Schlesinger, D.; de Souza Moura, P. N.

2026-06-07 bioinformatics 10.64898/2026.06.05.729843 medRxiv
Top 0.1%
4.0%
Show abstract

Distinguishing loss-of-function (LOF) from gain-of-function (GOF) effects of missense variants is fundamental to understanding disease mechanisms and guiding therapeutic strategy, yet no large-scale, expert-curated benchmark has been publicly available for this task. Here we present GLOF (Gain and Loss Of Function), a dataset of 112,399 missense variants across 2,809 human genes, each classified as LOF, GOF, or neutral by board-certified clinical geneticists following ACMG guidelines. Pathogenic variants were sourced from ClinVar and annotated with their functional mechanism based on published functional studies, phenotype correlations, and established gene-disease relationships. Neutral variants were drawn from gnomAD v3.1 and validated against v4.1 using stringent population frequency filters. The dataset spans diverse protein families, includes 97 genes with bidirectional mechanisms (containing both LOF and GOF variants), and has been validated against well-characterized variants in the literature. GLOF is publicly available on Kaggle (https://www.kaggle.com/datasets/maricatovictor/loss-and-gain-of-function-variants) and Hugging Face (https://huggingface.co/datasets/victormaricato/glof), and provides a standardized resource for developing and benchmarking computational methods that predict variant functional mechanisms.

19
First-Trimester Non-Invasive Prediction of Preterm Birth Using Cell-Free DNA Fragmentomics

Pham, M.-D. N.; Phan, M.-T. T.; Tran, N.-T.; Vo, T.-S.; Le, H.-T.; Nguyen, T.-H. T.; Nguyen, Q.-H. V.; Ha, M.-T. T.; Le, T. M.; Hoang, D.-T. T.; Huynh, K.-T. N.; Nguyen, N. V.; Nguyen, C. C.; Bui, T. C.; Nguyen, X. T.; Le, S. V.; Tran, V. D.; Nguyen, M.-N. B.; Nguyen, T. V.; Nguyen, T.-A. T.; Hoang, B. P.; Nguyen, T. V.; Nguyen, T.-A. T.; Nguyen, T. T.; Duong, T. D.; Pham, C. H.; Luong, K.-O. T.; Dao, C. N.; Hoang, K. V.; Huynh, T.-T. T.; Nguyen, K. M.; Tran, S.-T. T.; Tran, H. T.; Nguyen, S. C.; Tran, T. D.; Nguyen, P. T. L.; Pham, T. V.; Pham, K. C.; Thai, M. D.; Do, T.-T. T.; Dao, H. T.; Va

2026-07-11 genomics 10.64898/2026.07.07.736241 medRxiv
Top 0.2%
3.3%
Show abstract

ObjectiveTo develop and validate a cell-free DNA (cfDNA) fragmentomic classifier for the early prediction of spontaneous preterm birth (PTB) using routine first-trimester non-invasive prenatal testing (NIPT) data. MethodsA nested case-control study was conducted within a prospective multicenter Vietnamese cohort comprising 286 pregnancies, including 82 spontaneous PTB cases and 204 term controls. Maternal plasma cfDNA collected during routine first-trimester NIPT (median gestational age, 12 weeks) was sequenced to a depth of approximately 20 million reads per sample. Five fragmentomic feature categories including copy number alterations, end-motif composition, nucleosome distance, fragment length, and joint fragment-lengthxend-motif were evaluated for PTB prediction. Machine learning classifiers were developed in a training cohort (n = 228, 65 PTB vs 163TB) and tested in a validation cohort (n = 58, 17 PTB vs 41 TB). ResultsAmong the five fragmentomic feature classes evaluated, 4-mer end-motif (EM) profiles exhibited the most pronounced differences between PTB and term control samples. Consistent with these findings, the EM-based classifier demonstrated the highest discriminative performance in the validation cohort, achieving an AUC of 0.970 (95% CI, 0.912-1.000). At a specificity >90%, the model achieved a sensitivity of 94% (95% CI, 78-100%). ConclusionThese findings demonstrate that cfDNA EM signatures derived from routine first-trimester NIPT can accurately identify pregnancies at risk of spontaneous preterm birth, without additional blood collection or sequencing, thereby extending the clinical utility of existing prenatal screening infrastructure. KEY POINTSO_ST_ABSWhat is already known about this topic?C_ST_ABSO_LICurrent first-trimester prediction strategies based on maternal characteristics, cervical length, and biochemical markers have limited predictive accuracy, particularly in nulliparous women. C_LIO_LIExisting cfDNA-based approaches have shown only modest performance or require additional assays, limiting clinical applicability. C_LI What does this study add?O_LIExisting NIPT sequencing data can be repurposed (without additional blood sampling or sequencing) for accurate prediction of spontaneous preterm birth (AUC=0.970). C_LIO_LIA classifier employing 4-mer end-motif (EM) profiles achieved an AUC of 0.970. At a specificity >90%, the model achieved a sensitivity of 94%. C_LI

20
Variantscape: Large Language Model-Driven Mining of Biomedical Literature for Clinical Interpretation of Cancer Variants

Wosny, M.; Blindu, A. S.; Boesch, M.; Peres, T.; Niederhauser, T.; Fruh, M.; Rothermundt, C.; Hastings, J.

2026-08-03 health informatics 10.64898/2026.08.02.26359492 medRxiv
Top 0.2%
3.3%
Show abstract

Background: Precision oncology relies on accurate interpretation of tumour-detected gene variants, to guide personalized treatment decisions. However, accurate interpretation of variants in context requires extensive information that is often buried within unstructured biomedical literature and obscured by inconsistent nomenclature, making manual retrieval labour-intensive and prone to omissions. Methods: To address this challenge, we developed Variantscape, a large-scale, automated pipeline and open-access web tool. It integrates traditional natural language processing methods with state-of-the-art large language models to extract, standardize, and analyze co-associations between genetic variants, cancer types, and therapeutic interventions from published biomedical abstracts. Findings: From over 3 million abstracts screened, 335,817 gene name-containing articles were eligible for downstream extraction. Among these, 7,423 (2.2%) simultaneously mentioned a variant, cancer type, and therapeutic agent, encompassing 3,902 unique variants across 98 cancer types and 388 therapeutic agents. This highlights the inefficiency of manual literature retrieval in molecular tumour board (MTB) workflows. Network analysis revealed 14,831 statistically significant co-associations, represented in a literature-derived graph with 4,388 nodes and 46,943 edges. Canonical alterations in well-studied cancers (e.g., BRAF V600E in melanoma) were strongly linked to established treatments, while several rare variants also emerged with high-confidence literature support. Interpretation: By applying large language models to biomedical literature, Variantscape enables scalable, context-aware extraction of trilateral variant-treatment-cancer relationships. This approach supports early evidence synthesis/hypothesis generation, highlights underrecognized or rare associations, and offers a practical resource for accelerating discovery and supporting precision oncology research and translation. Unlike static databases, Variantscape is continuously updatable and leverages large language model-based inference to uncover putative associations without manual curation. Variantscape has the potential to support MTB workflows and translational research by rapidly revealing signals from underlying abstracts.